Customer churn remains a major challenge in the telecommunications industry, where retaining existing customers is signi7cantly more cost-e8ective than acquiring new ones. This study uses the publicly available IBM Telco Customer Churn dataset consisting of 5634 preprocessed records to develop a predictive model for customer churn classi7cation. The dataset was cleaned and prepared through preprocessing steps including handling missing values and encoding categorical variables to ensure suitability for machine learning analysis. Four classical machine learning models—decision tree, random forest, XGBoost, and logistic regression—were selected due to their proven e8ectiveness in classi7cation tasks, interpretability, and strong performance in structured tabular datasets commonly used in churn prediction studies. These models were evaluated using standard performance metrics, including accuracy, precision, recall, F1-score, and ROC-AUC. To enhance predictive performance and stability, a tuned approach was applied. The random forest achieved an accuracy of 84.73%, precision of 85.20% for churn, recall of 84.07% for churn, F1-score of 84.62%, and ROC-AUC of 93.86%. The results demonstrate that combining well-tuned classical machine learning models can produce reliable and robust churn prediction performance. This study contributes by providing a systematic evaluation of classical models on a benchmark dataset with standardized preprocessing and comprehensive performance analysis for telecom churn prediction.
Loading....